Goto

Collaborating Authors

 forward progress


Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models

arXiv.org Artificial Intelligence

Legged robots are physically capable of navigating a diverse variety of environments and overcoming a wide range of obstructions. For example, in a search and rescue mission, a legged robot could climb over debris, crawl through gaps, and navigate out of dead ends. However, the robot's controller needs to respond intelligently to such varied obstacles, and this requires handling unexpected and unusual scenarios successfully. This presents an open challenge to current learning methods, which often struggle with generalization to the long tail of unexpected situations without heavy human supervision. To address this issue, we investigate how to leverage the broad knowledge about the structure of the world and commonsense reasoning capabilities of vision-language models (VLMs) to aid legged robots in handling difficult, ambiguous situations. We propose a system, VLM-Predictive Control (VLM-PC), combining two key components that we find to be crucial for eliciting on-the-fly, adaptive behavior selection with VLMs: (1) in-context adaptation over previous robot interactions and (2) planning multiple skills into the future and replanning. We evaluate VLM-PC on several challenging real-world obstacle courses, involving dead ends and climbing and crawling, on a Go1 quadruped robot. Our experiments show that by reasoning over the history of interactions and future plans, VLMs enable the robot to autonomously perceive, navigate, and act in a wide range of complex scenarios that would otherwise require environment-specific engineering or human guidance.


Value Function Decomposition for Iterative Design of Reinforcement Learning Agents

arXiv.org Artificial Intelligence

Designing reinforcement learning (RL) agents is typically a difficult process that requires numerous design iterations. Learning can fail for a multitude of reasons, and standard RL methods provide too few tools to provide insight into the exact cause. In this paper, we show how to integrate value decomposition into a broad class of actor-critic algorithms and use it to assist in the iterative agent-design process. Value decomposition separates a reward function into distinct components and learns value estimates for each. These value estimates provide insight into an agent's learning and decision-making process and enable new training methods to mitigate common problems. As a demonstration, we introduce SAC-D, a variant of soft actor-critic (SAC) adapted for value decomposition. SAC-D maintains similar performance to SAC, while learning a larger set of value predictions. We also introduce decomposition-based tools that exploit this information, including a new reward influence metric, which measures each reward component's effect on agent decision-making. Using these tools, we provide several demonstrations of decomposition's use in identifying and addressing problems in the design of both environments and agents. Value decomposition is broadly applicable and easy to incorporate into existing algorithms and workflows, making it a powerful tool in an RL practitioner's toolbox.


STELLAR Analytics for the Win – an A for Enterprise Analytics Mastery

#artificialintelligence

When I was in middle school (quite a few years ago), I started to realize that I was pretty good at math. I had done okay before, but the problems and concepts were becoming more difficult. Surprisingly, I was really "getting it", while some of my classmates encountered more challenges with the subject matter. The teacher motivated us with a special letter grade when our performance on a homework assignment or quiz was stellar -- a large letter grade "A" on the top of our assignment page, which she called a "big bold A". This grade was not simply recognition for getting a 90% score on the assignment, but was awarded for achieving 99% or 100%.


Pizza-making robot that can assemble and cook 300 pizzas every hour

Daily Mail - Science & tech

Not even your local pizza joint is safe from the forward progress of automation. At CES, a Seattle based Picnic showcased its automated pizza-making system that can swiftly assemble and cook pies with minimal human interaction. The system, which consists of three compact modular panels that assemble to form a conveyor belt, is capable of taking a pre-made pizza crust, adorning it with toppings, and cooking the pie to pre-specified doneness. What's even more compelling than the fact the pizza is made with little to no human input, however, is the speed at which Picnic's bot operates. According to CEO Clayton Wood, the bot can churn out an impressive 300 12-inch pizzas every hour when at max capacity.